Papers by Abhinav Sukumar Rao
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are widely used and engage millions of users from diverse contexts and cultures. |
| Approach: | They propose an evaluation framework to assess LLMs’ cultural adaptability by measuring their ability to judge social acceptability across varying levels of cultural norm specificity. |
| Outcome: | The proposed model shows stronger adaptability to English-centric cultures over those from the Global South. |
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to jailbreak large language models have been poorly studied . a recent study showed that non-expert users can jailbreak LLMs by manipulating their prompts . |
| Approach: | They propose a formalism and a taxonomy of known (and possible) jailbreaks . they propose generating a dataset of model outputs across 3700 jailbreak prompts a 'prompt' attack is a new attack popularly categorized as "prompting injection attacks" |
| Outcome: | The proposed model exploits 3700 jailbreak prompts over 4 tasks to analyze their effectiveness . authors show that the model can learn to perform a new task on unseen examples . |